Skip to content

Expose bundle-aware speculative defaults before model load - #3029

Merged
jjang-ai merged 35 commits into
mainfrom
review/qwen-spec-defaults-oct7
Oct 8, 2026
Merged

jjang-ai merged 35 commits into
mainfrom
review/qwen-spec-defaults-oct7

Conversation

@jjang-ai

@jjang-ai jjang-ai commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Selecting a supported Qwen bundle exposes speculative On (Adaptive) / Off (AR) before the model loads.

  • Flash-Next bundles with real native heads default to Adaptive.
  • Qwen 27B bundles that ship a compatible DFlash 2 drafter use it by default.
  • Explicit Off disables both native and external drafting. Reset restores the bundle default.
  • Migration changes only provenance-owned, untouched defaults; user choices are preserved.
  • Drafter weights are released on target unload unless another resident target shares the path.

Engine pin: merged osaurus-ai/vmlx-swift#565 at a468573fbe5b2c166d8d1dabea78c205c41e6716, updated at all six tracked sites. This includes the performance work, follow-up drafter cache repair, physical head validation, SSD n-gram defaults, residency accounting, and the final split-fusion/invalid-offset audit fixes.

App-side changes in this PR

  • Allocator scratch floor (492c66f8). Under Safe Auto the remaining-budget clamp could set MLX's freed-buffer pool to 0 for a large model, so every decode step re-allocated with a residency commit. Flash-Next 4S target forward went from 31.4 to 20.5 ms, live.
  • Reload on family-default transitions (f41de0a3). The family default loads without the MTP head for every family except Flash-Next, so moving between the family default and On (Adaptive) is now a load-input change. Before, a resident model could stay head-less and never speculate.
  • Tests brought up to date with the shipped engine contracts: the default mode is the family default; a selectable drafter needs complete weights; the picker anchor follows a renamed local.

Evidence

  • Live dev build, isolated profile, Safe Auto, plain agent, bundle sampler. One connected chat on 27B JANGH2:
    • before: 74 / 40 / 35 / 24 / 19 tok/s;
    • after: 110 / 50 / 70 / 37 / 36 tok/s (code, pasted image, follow-up, two essay follow-ups).
    • Reasoning-on + image and a reasoning-off replay also verified. All answers were correct.
  • Flash-Next family and Allosaurus: see the comments.
  • Local tests: the affected suites pass. HTTPHandlerEndpointTests.runtimeSettings_put_persistsAndReportsRuntimeEffects can fail when run in parallel with other suites that override the shared settings directory. It passes alone (3/3) and in CI.

Known behaviour, not a defect

Each reasoning on/off/effort setting keeps its own prefix-cache chain (it is part of the cache key), so the first turn after a change re-prefills.

No release is included. Final consuming-build audit evidence is recorded in the merge comments.

Eric added 21 commits October 6, 2026 07:04
Points the app at vmlx-swift 371f5c40: multi-row bit-exact BF16-affine
verify kernels, lane matmul with safe tiling and the DFlash2 width
chooser, native-MTP copy drafts, and the Qwen4 PLE page cache. Proof
build only; not for release.
Pins vmlx-swift perf/claude-swift-port-oct6 @ a4f99a73 (six pin sites).

Settings > Speculative Decoding now has three modes: Off (AR), Default
and On (Adaptive). Default is the engine's `.familyDefault`: Adaptive
native MTP for Qwen3.8 Flash-Next, and a Qwen 27B bundle's own dflash2/
drafter. A fresh install shows the chat picker's Native MTP row as On
for Flash-Next bundles, and the user can switch it off. The picker and
the phone snapshot show the per-bundle effective state.

ModelRuntime passes the bundle directory to drafter selection, so a
bundled dflash2/ drafter is found without a folder pick.

The Safe Auto materialized-load memory check is advisory: it logs a
warning instead of refusing. It refused Flash-Next JANG_4S on a 128 GB
Mac ("require ~92 GiB, only ~53 GiB available"); the model loads and
runs at 55-115 tok/s. Strict mode keeps its explicit refusal.
Pins vmlx-swift perf/claude-swift-port-oct6 @ 0e26bcb5 (six pin sites).

The engine now loads dense Qwen3.5 JANGH bundles (Qwen3.8-27B-JANGH2:
jangtq2 2-bit MLP banks). Measured in RunBench on max2: AR 30.6 tok/s
prose, DFlash2 28.6 / 57.7 / 135.0 / 141.0 tok/s (prose / code / easy
code / easy prose).
@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Current live qualification found a numerical blocker; this PR remains draft and unmerged.

Source tested: Osaurus ffb03f2, vmlx-swift fd9d925c, MLX fork5aba1efd. The later engine91726781 change is whitespace only; it is not claimed as a new app build.

Live isolated Release app, dense Qwen3.8-27B-JANGH2:

  • Visible composer prose:4117 prompt tokens,302 output tokens,23.0 tok/s; cold prefill12.237s (336.4 prompt tok/s),106 actual DFlash verify cycles, natural completion.
  • Visible follow-up:4507 prompt tokens,26 output tokens,19.2 tok/s, L2 disk hit1/stores2, natural completion.
  • These two UI rows are not a matched AR speed comparison. The attempted UI AR control delegated a helper and is excluded.

Controlled app HTTP diagnostic, same38-token prose, explicit greedy, fresh cache namespace, AR/Adaptive/Adaptive/AR:29.01 /29.39 /27.37 /26.44 tok/s. Clocks differed, so this does not establish a speedup. AR answers matched each other; Adaptive answers differed from AR and from each other.

Holding DFlash width at its trained8 produced identical outputs twice at27.61/27.46 tok/s, but still differed from AR. This is a diagnostic, not a proposed fixed-depth product change.

Causal teacher-forced test on the actual bundle: from the same prefilled snapshot, target row0 of widths5/8/16 differs from a one-token call before acceptance or rollback, both with and without LaneQMM and in eager/staged modes (12 comparisons, max absolute logit difference0.125). Lane installation also changes cold-prefill logits. Next work isolates the first differing operator and tests an arithmetic repair. Changing the controller alone is insufficient.

Local raw artifacts: /Users/eric/vmlx-private-evidence/qwen-spec-review-20261007/{app-live002.log,ui-dense-prose.png,ui-dense-followup.png,dense-prose-abba001,dense-fixed8-001,review/prefill-partition/dense001.json,review/DFLASH-DENSE-ROW-CAUSAL.md}. These local artifacts are not publicly accessible CI attachments. Other quants, complete media/cache/lifecycle coverage, and final app proof remain pending.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Correction to the preceding qualification comment: I applied the wrong numerical contract to dense27B DFlash2. The supplied Swift handoff READ-FIRST §5c explicitly documents NAX/Lane verification as speed-first and not bit-identical to AR. The strict AR row-exact gate applies to Flash-Next native MTP.

The measured S1-versus-wide differences are real, but they are not evidence of a new regression or a standalone merge blocker under that documented dense27B contract. No production kernel or adaptive-width change was made. The fixed8 run was diagnostic only.

Review resumes against the correct references: fast QMV vs its replaced S1 kernel; small NAX tile vs its replaced large tile; verification/acceptance and cache commit within the chosen execution path; default selection, live app speeds, cold prefill, tool continuation, and media fallback. Dense27B controlled prose27–29 tok/s is consistent with the handoff28.6 tok/s. The23 tok/s visible-composer run had4117 prompt tokens and is not a matched regression result.

The PR remains unmerged pending the remaining integration/CI proofs, not because dense DFlash differs from plain AR.

Eric added 4 commits October 7, 2026 04:25
A working-set estimate that meets the Safe Auto budget left a 0-byte
allocator headroom, so native-MTP and DFlash 2 generations ran with no
freed-buffer reuse: every decode step re-allocated its intermediates and
paid a Metal residency commit per buffer. The working set already prices
a >= 2 GiB scratch floor for exactly these buffers, so the remaining-budget
clamp may shrink the pool but never below that floor.

Measured live on Qwen3.8 Flash-Next JANG_4S under Safe Auto: 31.4 ms per
target forward with the starved pool vs 21.6 ms with a working pool.
@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Live app proof — osaurus dev build pinned to this engine head

Setup.

  • osaurus review/qwen-spec-defaults-oct7 (53231ce9), pinned to vmlx-swift a5a0689c (docs-only over the proven e40f9fa4).
  • Isolated profile, default Safe Auto memory profile.
  • Plain chat agent (tools and memory off), effort None, bundle sampler defaults.
  • Real composer: three connected turns per bundle (code → pasted image → text follow-up).
  • Speed = wall decode tok/s from the app's step log. Speculation counters come from the engine log.
Bundle T1 code T2 image (answer) T3 follow-up after image (TTFT) Speculation
Qwen3.8-27B JANG_4D 103.0 44.3 (red 7 ✓) 61.3 (0.89 s, 1,202 / 1,407 restored) DFlash 2 trees on all turns
Qwen3.8-27B JANGH2 74.1 40.5 (blue 4 ✓) 35.2 (1.6 s, 801 / 995 restored) DFlash 2 trees on all turns
Flash-Next 4S 71.0 46.6 (green 2 ✓) fully cached (0.76 s) native MTP, 2.1–3.8 / verify
Flash-Next 4M 52.8 56.0 (red 7 ✓) 57.0 (0.65 s) native MTP
Flash-Next 2L 39.0 (cold) 52.2 (blue 4 ✓) 72.0 (0.73 s) native MTP
Flash-Next 6S 37.0 (cold) 39.7 (green 2 ✓) 75.1 (0.84 s) native MTP
Flash-Next 1L 42.6 (cold) 52.9 (red 7) 58.9 (0.64 s) AR (no head, by design)
Allosaurus JANGH2 43.4 (cold) 46.0 (blue 4) 83.8 (0.64 s) native MTP
Flash-Next JANGH4 47.7 (cold) 43.1 (green 2 ✓) 69.7 (1.1 s) native MTP

AR references on the same machine: 27B 4D ~23, 27B JANGH2 ~29, Flash-Next ~50–55 tok/s. "Cold" = the first request
after loading a large bundle (expert pages and kernels warming). Warm turns on the same bundles run 52–84.

App-side defect found and fixed in osaurus (492c66f8)

Under Safe Auto, the remaining-budget clamp could set MLX's freed-buffer pool to 0 bytes for a large model. Every decode
step then re-allocated its intermediates, with a Metal residency commit per buffer.

  • Flash-Next 4S: 31.4 → 20.5 ms per target forward, live, same build and prompt.
  • The clamp now keeps the scratch floor the working-set estimate already admits.

Open, not claimed

  • 27B JANGH2 multi-row verify is ~3× its 1-row cost. The codebook MLP falls to the generic tile above one row, so
    prose speedup on that bundle is small. A dedicated verify kernel is in progress.
  • The first request after a large-bundle load is slow.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Pin → vmlx-swift e54e7388 (4c9e5e4e)

Two engine fixes for Qwen3.8-27B DFlash 2, re-proven live in this dev build. Details and tables are on vmlx-swift
PR #565.

  • JANGH dense verify tile (90eecf3c). 27B JANGH2 multi-row verify went from ~2x to ~1.55x a decode step.
  • Follow-up turns after a prefix restore (e54e7388). The drafter context was corrupted on every restored turn, so
    chat follow-ups on 27B ran below AR. Fixed.

Live, one connected chat on 27B JANGH2 (code → pasted image → follow-up → two essay follow-ups):

  • before: 74.1 / 40.5 / 35.2 / 23.7 / 18.7 tok/s
  • after: 109.8 / 49.5 / 70.2 / 37.2 / 36.1 tok/s

All answers were correct and coherent. No osaurus code changed beyond the pin.

For the audit: the restored follow-up turn is the shape to check for any speculative change. Single fresh-prompt
benchmarks did not show this defect.

The family default loads without the native MTP head for every family
except Qwen3.8 Flash-Next. Switching a resident model from the family
default to On (Adaptive) compared only off/not-off, so no reload ran and
the head-less model never speculated until a manual reload. A change into
or out of the family default is now a load-input change
(ServerControllerConfigLoadingTests already expected this and failed in CI).

Tests brought up to date with the shipped engine contracts:
- the settings default is the bundle-aware family default, not Off;
- a selectable external drafter needs complete weights (engine validates
  safetensors headers against the required shapes), so the fixture writes a
  sparse but complete artifact instead of a config-only folder;
- the picker blocked-bundle check follows the renamed local in
  FloatingInputCard (behaviour unchanged: blocked bundles render Off).
@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Ready for final audit at f41de0a3. Fixes the three CI test-core suites. One was a real defect: moving between the family default and On (Adaptive) did not reload the model, so a resident model could stay without its MTP head. The other two were stale test expectations. Engine pin e54e7388 is runtime-identical to the #565 head. Re-pin to the #565 merge commit at the six sites before merging. No release.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Independent audit update at app f41de0a3 / engine pin e54e7388 (runtime-identical to engine PR head 8e13b426). No runtime changes during this audit; merge is still held for the remaining performance checks. Fresh isolated Release app built successfully; all nine app CI checks pass.

Live app proof on Qwen3.8-27B-JANGH2:

  • Speculation On was visible before load. A real On chat ran DFlash2. Selecting Off then running another chat used ordinary decode: 47 output tokens, 27.7 engine tok/s, natural stop. On reactivation/persistence proof remains open.
  • Two serial file_read calls, followed by a second user turn with two more reads, exercised SSD persistence before each continuation. Published/restored boundaries progressed 4448 → 4663 → 4864 → 5032 → 5247 → 5448; restored DFlash2 physical offsets matched each boundary. No paged-RAM hits. Cold 4455-token prefill took 6288 ms; subsequent cached prompt stages took 455–610 ms. Cached full-prompt/time ratios are not cold-prefill throughput.
  • Correct verification-code/file extraction on the follow-up. The first answer made an arithmetic error (17+26+9 reported as44); this is retained as a model-quality failure, not hidden by the runtime/cache pass.
  • Physical footprint peaked at13.2 GiB and fell to524 MiB after unload. Legacy resolved prefix.memoryPercent=15 is not consumed by CacheCoordinator allocation; it does not establish a19-GiB RAM-prefix tier.

Remaining: speed comparison is not cleared. Current sampled full-matrix JANGH2 medians are35.36 prose/93.29 code/149.91 easy-code tok/s, below parts of the historical table; additional runs show clock and adaptive-width variation. No sampler/kernel changes have been made to chase these numbers. Isolated profile also exposed history-save errors in code unchanged from main, and duplicate bundle basenames make the short /v1/models alias ambiguous; the fully qualified bundle ID routes successfully. Neither is being silently called fixed.

Evidence receipts: app-build-current.log, app-live.log, tool-turn-debug.log, ram-recheck-vmmap.txt, h2-speed-current.jsonl, gpu.log, and CURRENT-AUDIT.md in the private PR565/PR3029 audit folder. No release/tag/appcast work.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Additional independent live app audit on f41de0a / engine e54e7388:

  • Off persisted across app restart. Switching On produced 21 real DFlash verify cycles and correct Python print(1)..print(50), natural stop: 295 generated tokens at 130.7 engine tok/s. This prompt includes history and seven tools; it is not the bare 100-print benchmark.
  • Attached a synthetic red image with a white 7 through the app file picker. Initial and follow-up answers both correctly identified the digit and colors, completed naturally, and restored the input controls.
  • Initial media request published SSD prefix 2113. Follow-up restored 2113 with 114 suffix tokens; drafter physical KV length and offset were both 2113. Follow-up then published prefix 2220. Media salt stayed identical. Prefill stage fell from 2746 ms (2120-token cold media prompt) to 274 ms (2227-token cached follow-up).
  • Engine generation rates for these short 14-token visual answers were 26.4 and 47.8 tok/s; these are cache/correctness probes, not sustained throughput benchmarks.
  • Physical footprint peaked at 11.1 GiB and returned to 501 MiB after unload. Paged RAM cache remained off. The legacy prefix-memory-percentage field does not allocate a RAM tier in this configuration.

Found one misleading baseline diagnostic: nil nativeMTPStats logs plain/off even for actual DFlash2 cycles. A diagnostics-only correction is under build verification; no decoder/kernel/cache/sampler changes.

Merge remains held: current H2 sustained prose speed is not cleared against historical receipts, and remaining quant coverage is incomplete. Raw private receipts: pr565-pr3029-audit-20261007/app-live-restarted.log, app-media-prefill-debug.log, media-cache-stats.json, media-physical-footprint.txt.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Diagnostics correction pushed as 9167afe: nil native-MTP stats no longer claim plain AR. The line now states nativeMTP=off decodePath=unreported, leaving actual DFlash path proof to engine telemetry. No generation, cache or kernel behavior changed.

Fresh isolated Release build passed. In the rebuilt app, attached a neutral-named image with yellow2/blue background after the previous white7/red image. Both the initial answer and follow-up were correct and completed naturally. Changed media used a different salt; follow-up restored2340 tokens with111 suffix tokens and aligned drafter offsets. New diagnostic text is observed live alongside real DFlash cycles. App history survived relaunch.

These short correctness rows overlapped CPU compilation and are not speed gates:15tokens22.6tok/s and8tokens13.9tok/s. Remaining sustained-speed audit and quant coverage are still open; not merge-ready yet.

Receipts: app-telemetry-build.log, app-telemetry-live.log, app-telemetry-prefill-debug.log in private pr565-pr3029-audit-20261007 evidence.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Independent current-engine checks (still partial; not a merge-ready claim):

  • Release app rebuilt with engine dc2c439c, signed and relaunched in its isolated profile with family_default. App picker shows Adaptive before load for 27B4D/H2, Flash2L/4S/4M/6S/H4 and AllosaurusH2. Headless Flash1L exposes no speculative row. CUA AX plus screenshots inspected.
  • Production metadata resolver executed 26 decisions: all nine installed bundles plus renamed/headless/missing-drafter views under Default and Off. Renaming preserves capability. All Off rows select AR. No weights loaded in this resolver check.
  • Actual optional-artifact controls: original bundles unchanged. A private 27BH2 view omitting only dflash2 loads Qwen35 under Default/AR and completes579/662-token turns; headless Flash1L loads Qwen4Exp under Default/AR and completes574/494-token turns. Both second turns restore SSD prefix28; this is partial-prefix reuse, not complete generated-answer reuse. No DFlash/native verify cycles in either control.
  • Observed rates: drafter-free27BH2 21.97/21.51tok/s; headless1L 30.97/29.85tok/s. CPU builds overlapped: these are load/cache-correctness probes, not valid speed comparisons. Their harness retained only answer previews, so they do not replace full visible-answer UI rows.
  • The packaged eval executable is building; delegated actual strategy and remaining per-quant speed/cold-prefill/UI rows are still open. Nothing merged or released.

SOURCE EVIDENCE: engine resolver/current headers, app bundle-capability picker path; bundle-resolution-receipt.json records binarySHA and exact paths.
LIVE EVIDENCE: app-engine-pin-build.log/app-dc2-identity.json; current CUA picker captures; no-drafter-hardlink-live.log and headless-1l-live.log plus receipts. Local audit folder: ~/vmlx-swift/docs/internal/pr565-pr3029-audit-20261007/.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Pushed a96dae9f: the MiMo/N2 launch exemption now requires matching bundle architecture as well as the legacy name. A Qwen bundle renamed to a MiMo/N2-looking alias continues through actual head inspection. Explicit Off remains authoritative. True MiMo architecture controls remain covered.

129 focused tests passed across NativeMTPPreloadDetectionTests, NativeMTPAdmissionTests and RuntimePolicySourceTests before the mechanical engine pin update. Added adversarial Qwen dense/Flash affine/JANGH aliases, malformed/missing config, conflicting nested architecture, and explicit Off assertions.

All six tracked runtime pin sites now point to engine 59f9323c, whose cache-tail correction and repeated real Allosaurus proof are recorded in osaurus-ai/vmlx-swift#565 (comment) . Exact repinned Release app is rebuilding. Do not treat the preceding app tests as runtime proof of the new pin.

SOURCE EVIDENCE: ModelFamilyNames.swift, ModelRuntime.swift and NativeMTPPreloadDetectionTests.swift at a96dae9f.
LIVE EVIDENCE: app-bundle-alias-tests.log (129/129); corrected-app model/UI proof remains pending. Existing source-resolution checks are not an all-bundle speed pass. Both PRs remain unmerged.

Eric added 3 commits October 7, 2026 15:34
The subagent runner always set prompt_tokens from the message estimator,
which omits the rendered chat template and the separately attached tool
schemas. A 4-tool delegated child reported 92 prompt tokens while the
runtime prefilled 1,455 (same 1,427 completion tokens / 48.5 tok/s), and
total, worker and context-saved figures inherited the gap.

The runner now takes the count from the input-token hint the runtime
already streams (or the stats sentinel's input count) and keeps the
estimator only as the fallback when neither arrives. Completion and
tokens-per-second are unchanged.

Tests: runtime count wins over the estimate; estimator fallback when no
count is streamed.
@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Audit follow-up — head 13839484, pinned to vmlx-swift 7c5a8a08

  • d546b3a0 Delegated usage. Reports the runtime's prompt tokens (input-token hint / stats input count); the
    estimator is now only a fallback. It was 92 reported vs 1,455 prefilled.
  • f6700d3b, 13839484 Re-pins at all six sites.
  • Tests. 258 tests across RuntimePolicySource, ImageGenerationBridgeContract, NativeMTP*, ServerControllerConfigLoading,
    SpawnTool, ModelRuntimeRAMFeasibility and ServerRuntimeSettingsStore pass at this pin.

Re-pin to the vmlx-swift #565 merge SHA before merging. No release.

@jjang-ai

jjang-ai commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Final consuming audit: app 15e917f pins merged engine a468573fbe5b2c166d8d1dabea78c205c41e6716 at all six tracked sites. Engine #565 is merged. The merge tree exactly matches tested engine f3963efb7.

SOURCE EVIDENCE: the final app delta is only those six pin replacements; all existing performance, mode-reload, drafter discovery, allocator scratch-floor and delegation usage changes remain in this PR. The 30-second inactivity setting/toggle is unchanged. Engine final audit adds only split-fusion completeness and bounds-safe residency metadata accounting; it does not change decode kernels, sampling, cache storage or adaptive scheduling.

LIVE EVIDENCE / TESTS:

  • Consuming swift build --build-system swiftbuild --build-tests passed, 220.47 s; dependency checkout verified at a468573f.
  • Consuming affected suites: 258 tests / 15 suites passed. Filter: RuntimePolicySourceTests|ImageGenerationBridgeContractTests|NativeMTP|ServerControllerConfigLoadingTests|SpawnToolTests|ModelRuntimeRAMFeasibilityTests|ServerRuntimeSettingsStoreTests.
  • Final engine resolver: all nine real bundles preserve routing. Default selects native MTP for Flash 2L/4S/4M/6S/JANGH4 and Allosaurus JANGH2, DFlash2 for shipped 27B 4D/JANGH2 drafter bundles, and AR for headless 1L. Explicit Off suppresses every route.
  • Retained unchanged live UI, per-tool SSD, media, Default/Off delegation/eval and per-quant timing evidence from prior comments and the private handoff remains applicable; no redundant benchmark campaign was run for a six-pin app delta.
  • Independently inspected latest Allosaurus app follow-up: 9,230 prompt / 596 generated, normal stop; 0.976 s TTFT; 101.4 decode tok/s versus 98.3 runtime generation tok/s; 103 verifies, 5.73 committed/verify, .996 acceptance. SSD diskL2Hits 2→3, stores 3→5. App footprint 11–13 GB; system wired/file-cache counters are not app-owned memory.

Raw current artifacts: private merge-final-delta-20261007/app-build.log, app-tests.log, real-bundles-{family_default,off}.log and receipt. Frozen prior live evidence is indexed under pr565-pr3029-audit-20261007/claude-live-consolidation/README.md. The bounded earlier 27B H2 Off delegation timeout is not represented as a completed-answer pass. Restored TTFT is not cold-prefill throughput.

GitHub CI is separately reported by its checks: running/queued checks are not claimed green. Local affected tests above passed at the exact consuming pin; engine repository-wide formatter drift and inherited quant-metadata tests remain documented. No release, tag, appcast, model-unload policy change or new performance tuning is included.

@jjang-ai
jjang-ai marked this pull request as ready for review October 7, 2026 23:56
@jjang-ai
jjang-ai merged commit 3e596da into main Oct 8, 2026
9 checks passed
@jjang-ai
jjang-ai deleted the review/qwen-spec-defaults-oct7 branch October 8, 2026 00:42
@jjang-ai

jjang-ai commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Merged after all app CI checks passed: https://github.com/osaurus-ai/osaurus/actions/runs/37704672816.

Osaurus main is 3e596da006db13d1bbf1a996b3da2a101283c2e7, consuming merged engine a468573fbe5b2c166d8d1dabea78c205c41e6716 at all six sites. The app merge tree matches the tested head exactly; the engine merge tree also matches its tested head. The repository-required squash merge was used; no admin bypass.

Local proof: 258 affected app tests, 6 engine delta tests, 9 real bundle Default/Off resolver pairs. Existing live speed/cache/media/delegation evidence is retained with the limits in the preceding audit comments. 30-second inactivity setting unchanged. No release/tag/appcast. Internal final handoff and hashed stop receipt are saved; work stopped.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant